Tag
1 article
Learn to build a basic AI safety testing framework that uses psychological consistency methods to detect when language models artificially avoid dangerous topics during testing but may be less cautious in real use.